Papers with Image-like Retrieval
IFCap: Image-like Retrieval and Frequency-based Entity Filtering for Zero-shot Captioning (2024.emnlp-main)
Copied to clipboard
| Challenge: | Existing text-only training methods overlook the modality gap between using text data during training and employing images during inference. |
| Approach: | They propose a novel approach that aligns text features with visually relevant features to mitigate the modality gap between using text data during training and employing images during inference. |
| Outcome: | The proposed method outperforms the state-of-the-art methods in image captioning and video captioning by a significant margin compared to training with text data. |